Tested on M1 Max: Diffusers stretches an 832×1216 reference to 832×1248, making figures 2.7% narrower. Setting output_resolution=1006 cut whole-image drift to about 0.1px.
Tested on M1 Max 64GB: the training-free TaylorSeer cache in Diffusers made Qwen-Image 2.1 about 2.7x faster. It crashes with the default KV cache; a small patch fixes it.
Tested Qwen-Image 2.1 open weights on M1 Max 64GB: overhead shots work locally, ~10 min per 832×1216 image (2x+ Anima), and i2i kept the face. Same prompts as my Anima/WAI tests.
Tested on M1 Max 64GB: AtomicChat's M64 GGUF keeps the 51B N-gram table in its own shard, so a llama.cpp PR #27742 build leaves it on SSD and runs the 125B MoE at 17.6 tok/s.
Tested on M1 Max 64GB with MLX: LLaMA Pro-style block expansion on Qwen3-0.6B-Base vs LoRA vs full FT, trained on 722 blog posts. Held-out PPL drops 14.7 to 10.5, but Wikipedia-ja rises 13.5 to 17.9, more forgetting than full FT. A 20% Wikipedia mix nearly removes it, and renumbering 28-layer LoRA keys reproduces the Anima-2.9B result in numbers.
Benchmarked on an M4 Mac mini: Ben Joffe's 2-instruction weekday hack beats plain %7 by 1.6-6.4x in clang, Rust, and V8, loses 3x in CPython. Plus a 25x V8 -0 deopt trap.
Tested on M1 Max 64GB: Qwen3.8-27B hits ~19 tok/s on both MLX and Ollama, but the default reasoning_effort=xhigh blew thinking up to 50,373 chars. Why Ollama dodges it.
Tested on M1 Max 64GB: hooked Qwen3.6-35B-A3B's MoE router in mlx-lm, pre-warmed the top-20 hot experts, still ~62 tok/s vs plain mmap cache. Plus the Metal OOM on Qwen3.5-122B.
Tested on M1 Max 64GB: SeFi-Image turbo runs on MPS in bf16 at 13–47s/image; fake text and face artifacts only clear up on 5B RL at 50 steps, 18 min/image.
Measured on M1 Max 64GB: weight-only INT8 runs 41% slower than fp16, a hand-written Metal int8 GEMM 15.6x slower, MPSMatrix int8 3x, and MLX 8-bit 9% slower.
Tested on M1 Max 64GB ComfyUI: v0.24.1 fails to load int8_tensorwise, v0.30.1 hits the missing aten::_int_mm MPS kernel, and a dequantize patch runs slower than bf16.
Tested on M1 Max 64GB: AnimaLoraToolkit + Anima-Base crashes on torch 2.6.0 (MPS SDPA bug) but runs 300 steps clean on 2.7.1, at ~21 s/step, roughly 9x slower than an RTX 5090.